Papers with Cross-modal Attention Congruence Regularization
Cross-modal Attention Congruence Regularization for Vision-Language Relation Alignment (2023.acl-long)
Copied to clipboard
| Challenge: | Despite recent progress towards scaling up multimodal vision-language models, these models struggle on compositional generalization benchmarks such as Winoground. |
| Approach: | They propose to use a cross-modal attention regularization loss to enforce relation alignment by capturing the semantic relation ‘in’ to match the visual attention from the mug to the grass. |
| Outcome: | The proposed approach improves Winoground Group score by 5.75 points . |